Generates videos from text, images, video, and audio references, with optional audio generation. Supports videos up to 30 seconds with customizable resolution, aspect ratio, and output format.
OPEN HIGGSFIELD API
Sign up for 15% off all sale models - then raise your favorites up to 50%, plus $15 on your balance. Your first result is minutes away
Integration
CONNECT ONCE. BUILD ANYTHING.
1const res = await fetch(2 'https://api.higgsfield.ai/bytedance/seedance-2.0/text-to-video', {3 method: 'POST',4 headers: {5 Authorization: `Key ${HF_API_KEY_ID}:${HF_API_KEY_SECRET}`,6 'Content-Type': 'application/json',7 },8 body: JSON.stringify({9 prompt: 'A cinematic tracking shot along a sunlit coastal road',10 resolution: '720p',11 generate_audio: true,12 duration: 5,13 aspect_ratio: '16:9',14 })15 }16)- Seedance 2.5
- Genjutsu
- Kling 3.0
- MiniMax H3
- Wan 3.0 Prime
- Marketing Studio Image
- Seedance 2.0
- Cinema Studio 4.0
- Grok Imagine 2.0
- Soul 2
- Genjutsu
- Wan 3.0
- LTX 2.5 Fast
- LTX 2.5 Pro
- Grok Imagine Video 1.5
- Wan 2.7
- Happy Horse 1.1
- Ideogram 4.0
- Recraft 4.1
- Happy Horse 1.0
- PixVerse 6
- Kling O3
- Kling 2.6
- Kling O1 (Omni)
- Wan 2.6
- MiniMax Hailuo 2.3
- Kling 2.5
- Soul Standard
- Cinema Studio 4.0
- Qwen Image 3
- Z-Image Turbo
- Product shots
- Graphic ads
- Marketplace design
Model catalog
THE MODELS YOU NEED, ALL IN ONE PLACE.
Transforms existing videos using image references for different characters, locations, and styles, preserving the original motion, camera movement, and timing throughout the resulting video clip.
Generates multi-shot videos from text or images, with native audio and durations up to 15 seconds.
Generates 2K videos up to 15 seconds using text, image, video, and optional audio reference files.
Generates videos from text or media references with fast processing and clips up to 30 seconds.

Generates and edits campaign images from text and image inputs, with resolutions from 1K to 4K. Preset mode uses a product photo and an optional model reference to guide the resulting image.
Generates videos from text, images, video, and audio references, with optional audio generation. Supports videos up to 30 seconds with customizable resolution, aspect ratio, and output format.
Transforms existing videos using image references for different characters, locations, and styles, preserving the original motion, camera movement, and timing throughout the resulting video clip.
Generates multi-shot videos from text or images, with native audio and durations up to 15 seconds.
Generates 2K videos up to 15 seconds using text, image, video, and optional audio reference files.
Generates videos from text or media references with fast processing and clips up to 30 seconds.

Generates and edits campaign images from text and image inputs, with resolutions from 1K to 4K. Preset mode uses a product photo and an optional model reference to guide the resulting image.
Creates 4K video from text or media references. Supports up to 15 seconds with optional sound.
Generates cinematic videos from text, image, or video references with automatic scene direction. Supports videos up to 30 seconds, optional sound, and customizable resolution and aspect ratio.

Generates images and makes precise edits while preserving image details and visual consistency.

Generates realistic portraits and fashion images, with natural textures and curated editorial styles.
Transfer motion from a reference video to your characters, products, or clothes.
Generates videos from text or images with native audio and adjustable duration up to 30 seconds.
Generates videos from text or images with native audio, camera motion controls and 4K resolution.
Generates videos from text or images with native audio, camera motion controls and 1080p output.
Creates videos from text prompts or still images, with 1080p resolution and clips up to 15 seconds.
Generates videos from text or images with native audio and consistent characters across scenes.
Generates videos from text or image references with output at 1080p and clips up to 15 seconds.

Generates and edits images with multilingual text for posters, packaging and other graphic layouts.

Generates images and illustrations using custom palettes and background colors to fit your brand.
Generates videos from text or image references with clips up to 15 seconds and 1080p resolution.
Generates videos from text or images with native sound and control over the first and last frames.
Generates and edits videos using image or video references, multi-shot scenes and native sound.
Generates videos and sound from text or images, including dialogue, sound effects and ambience.
Generates and edits videos using image or video references and controls for start and end frames.
Generates multi-shot videos from text or images with native sound and optional video references.
Generates short videos from text or images with fluid body motion and natural facial expressions.
Generates videos from text prompts and images, with fast generation and precise prompt control.

Generates fashion and lifestyle photos from text, with style presets for editorial and social content.
Cinematic video generation with automatic scene direction and optional references

Generates images from English or Chinese text, with enhanced prompts and resolution up to 2K.

Generates images quickly with clear Chinese and English text rendering and PNG output up to 2K.

Turns product photos into campaign images with presets for composition, lighting and visual style.

Creates ad creatives from a product photo, with presets that guide layout, typography and colors.

Creates listing images from a product photo, with preset styles designed for marketplace catalogs.
7-DAY LAUNCH OFFER
PICK ANY 4 MODELS. SAVE UP TO 50%
- Seedance 2.530% off$0.144 / sec$0.2057
- Genjutsu50% off$0.159 / sec$0.318
- Kling 3.045% off$0.0462 / sec$0.084
- MiniMax H330% off$0.091 / sec$0.13
- Wan 3.0 Prime30% off$0.0476 / sec$0.068
- Marketing Studio Image$0.0126 / image
- Seedance 2.030% off$0.0985 / sec$0.1407
- Cinema Studio 4.0$0.2057 / sec
- Grok Imagine 2.0$0.04 / image
- Soul 2$0.0032 / image
- Genjutsu50% off$0.159 / sec$0.318
- Wan 3.050% off$0.025 / sec$0.05
- LTX 2.5 Fast$0.09 / sec
- LTX 2.5 Pro$0.12 / sec
- Grok Imagine Video 1.5$0.08 / sec
- Wan 2.750% off$0.05 / sec$0.10
- Happy Horse 1.130% off$0.098 / sec$0.14
- Ideogram 4.0$0.03 / image
- Recraft 4.1$0.035 / image
- Happy Horse 1.030% off$0.098 / sec$0.14
- PixVerse 615% off$0.0978 / sec$0.115
- Kling O345% off$0.0462 / sec$0.084
- Kling 2.645% off$0.0385 / sec$0.07
- Kling O1 (Omni)45% off$0.0462 / sec$0.084
- Wan 2.650% off$0.05 / sec$0.10
- MiniMax Hailuo 2.375% off$0.0117 / sec$0.0467
- Kling 2.545% off$0.0231 / sec$0.042
- Soul Standard$0.0938 / image
- Cinema Studio 4.0
- Qwen Image 3$0.04 / image
- Z-Image Turbo$0.015 / image
- Product shots
- Graphic ads
- Marketplace design
Ideas to build
Don't start from a blank page
Content factory
Write one line about your product and get a creator-style video with the script, presenter, and voice ready.
Content factory Content factory Content factory Content factory Content factory Content factory
Fashion photoshoots
A full photoshoot without the studio. Consistent AI models wear your products in your brand's style - shot, retouched, and e-commerce ready in minutes, not weeks

Cinematic pipeline
Your idea, now in motion. Turn a short creative brief into a polished cinematic sequence without a production crew, expensive equipment, or weeks of post-production.
Works with your existing stack
Teams already using fal or Replicate can switch by swapping a single key and keep the rest of their code in place.
1 week offer
Sign up and save up to 50%
Generates videos from text, images, video, and audio references, with optional audio generation. Supports videos up to 30 seconds with customizable resolution, aspect ratio, and output format.
Transforms existing videos using image references for different characters, locations, and styles, preserving the original motion, camera movement, and timing throughout the resulting video clip.
Generates multi-shot videos from text or images, with native audio and durations up to 15 seconds.
Generates 2K videos up to 15 seconds using text, image, video, and optional audio reference files.
Generates videos from text or media references with fast processing and clips up to 30 seconds.

Generates and edits campaign images from text and image inputs, with resolutions from 1K to 4K. Preset mode uses a product photo and an optional model reference to guide the resulting image.
Creates 4K video from text or media references. Supports up to 15 seconds with optional sound.
Generates cinematic videos from text, image, or video references with automatic scene direction. Supports videos up to 30 seconds, optional sound, and customizable resolution and aspect ratio.
Analytics
See where every dollar goes
Track spend by day, model, and workflow, with discounts shown on your invoice. Pick a period, export the CSV, and send it to finance.
Get started
Three steps from key to content
The same flow for every image, video, and audio model. JavaScript, Python, and cURL examples are all in the docs.
- 1
Configure credentials
export HF_API_KEY_ID="your-api-key-id"export HF_API_KEY_SECRET="your-api-key-secret" - 2
Submit a generation
curl --request POST \ --url https://api.higgsfield.ai/bytedance/seedance-2.0/text-to-video \ --header "Authorization: Key ${HF_API_KEY_ID}:${HF_API_KEY_SECRET}" \ --header "Content-Type: application/json" \ --data '{ "prompt": "A cinematic tracking shot along a sunlit coastal road", "resolution": "720p", "generate_audio": true, "duration": 5, "aspect_ratio": "16:9" }' - 3
Check the result
curl --silent --show-error --fail-with-body \ --url "https://api.higgsfield.ai/requests/${REQUEST_ID}/status" \ --header "Authorization: Key ${HF_API_KEY_ID}:${HF_API_KEY_SECRET}" | jq
Community over 25 MILLION USERS
Join a global creative network where people generate AI images, share ideas, and inspire each other every day.




Trusted by 5.000+ people worldwide
Still have questions?
We’ve answered the most frequently asked questions

